Papers by Junyi Jessy Li

53 papers
Detection and Measurement of Syntactic Templates in Generated Text (2024.emnlp-main)

Copied to clipboard

Challenge: Existing diversity evaluation focuses primarily on word-level features.
Approach: They propose a method for evaluating diversity over syntactic features to characterize general repetition in large language models.
Outcome: The proposed method shows that models produce templated text in downstream tasks at a higher rate than what is found in human-reference texts.
QUDeval: The Evaluation of Questions Under Discussion Discourse Parsing (2023.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics poorly approximate parser quality, says a new study . questions under discussion is a linguistic framework that views discourse as asking questions and answering them .
Approach: They propose a framework for automatic evaluation of QUD parsing . they use a dataset of fine-grained evaluation of 2,190 QUD questions .
Outcome: The proposed framework shows that satisfying constraints of QUD is still challenging for modern LLMs.
Training Dynamics for Text Summarization Models (2022.findings-acl)

Copied to clipboard

Challenge: Pre-trained language models have shown impressive results when fine-tuned on large summarization datasets.
Approach: They analyze the training dynamics for generation models, focusing on summarization . they find that a propensity to copy the input is learned early in the training process .
Outcome: The proposed model learns at different stages of fine-tuning, the authors show . they show that factual errors are learnt in later stages, but not at high-loss tokens .
InfoLossQA: Characterizing and Recovering Information Loss in Text Simplification (2024.acl-long)

Copied to clipboard

Challenge: Text simplification aims to make technical texts more accessible to laypeople but often results in deletion of information and vagueness.
Approach: They propose a framework to characterize and recover simplification-induced information loss in form of question-and-answer (QA) pairs.
Outcome: The proposed framework characterizes and recovers simplification-induced information loss in form of question-and-answer (QA) pairs.
Multilingual Simplification of Medical Texts (2023.emnlp-main)

Copied to clipboard

Challenge: Existing work on medical text simplification has focused on monolingual settings . important findings in medicine are typically presented in technical, jargon-laden language . text simulating models can generate viable simplified texts, but there are outstanding challenges .
Approach: They propose a dataset for medical text simplification in four languages . they evaluate fine-tuned and zero-shot models across these languages based on human assessments and analyses .
Outcome: The proposed dataset evaluates models in English, Spanish, French, and Farsi . it shows that the models can generate viable simplified texts, but there are challenges .
Elaborative Simplification: Content Addition and Explanation Generation in Text Simplification (2021.findings-acl)

Copied to clipboard

Challenge: a new study examines the use of content addition in text simplification when complex concepts need to be explained.
Approach: They present a data-driven study of content addition in text simplification . they analyze 1.3K instances of elaborative simplification in the Newsela corpus .
Outcome: The proposed study shows that contextual specificity can improve elaboration generation performance.
Summarizing, Simplifying, and Synthesizing Medical Evidence using GPT-3 (with Varying Success) (2023.acl-short)

Copied to clipboard

Challenge: Large language models are capable of producing high quality summaries of general domain news articles in few- and zero-shot settings, but it is unclear whether they are similarly capable in more specialized domains such as biomedicine.
Approach: They use GPT-3 to generate single- and multi-document summaries of biomedical articles, given no supervision, using a set of annotations.
Outcome: The proposed model outperforms fully supervised models in generic news summarization, but struggles to synthesize evidence across multiple documents.
LinkNav: Surfacing Interconnected Information in Scientific Articles (2026.acl-demo)

Copied to clipboard

Challenge: a non-linear reading order of academic literature is recognized by authors who make explicit connections between non-adjacent passages.
Approach: They propose an enhanced reading experience which generates questions and searches for answer-bearing passages in academic papers to form intra-document connections when answers are found.
Outcome: The proposed interface makes connections between related but non-adjacent passages even if the author did not make them explicit.
Why Swear? Analyzing and Inferring the Intentions of Vulgar Expressions (D18-1)

Copied to clipboard

Challenge: Vulgar words are employed in language use for several different functions, including expressing aggression, signaling group identity or the informality of the communication.
Approach: They present a dataset of 7,800 tweets with six categories of vulgarity in which all instances of vulgar words are annotated with one of the six categories.
Outcome: The proposed model can predict the category of a vulgar word based on the immediate context it appears in with 67.4 macro F1 across six classes.
Evaluating Subjective Cognitive Appraisals of Emotions from Large Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing work on automatic prediction of cognitive appraisals has focused on physiological aspects of emotions.
Approach: They present a dataset that assesses 24 appraisal dimensions across 241 Reddit posts . they find that open-source models fail to automatically assess and explain cognitive appraisals .
Outcome: The proposed dataset assesses 24 appraisal dimensions across 241 reddit posts.
Learning to Update Natural Language Comments Based on Code Changes (2020.acl-main)

Copied to clipboard

Challenge: a novel approach to update comments based on code changes is proposed . a dataset of open-source software projects is used to train and evaluate the model .
Approach: They propose an approach that learns to correlate changes across two distinct language representations to generate a sequence of edits that are applied to the existing comment to reflect the source code modifications.
Outcome: The proposed model outperforms baselines and automatic metrics with respect to making edits.
Linguistically-Informed Specificity and Semantic Plausibility for Dialogue Generation (N19-1)

Copied to clipboard

Challenge: Past work has focused on word frequency-based approaches to improving specificity, such as penalizing responses with only common words.
Approach: They propose to rerank a sequence-to-sequence model to improve the informativeness, reasonableness, and grammatically of responses by using externally-trained classifiers targeting each of these factors.
Outcome: The proposed model improves the informativeness, reasonableness, and grammatically of responses.
Paragraph-level Simplification of Medical Texts (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods for simplification of medical texts are limited due to jargon and technical content.
Approach: They propose to automate the simplification of medical texts by penalizing decoders for producing "jargon" terms.
Outcome: The proposed method improves on existing heuristics by penalizing the decoder for producing "jargon" terms.
An Annotated Dataset of Discourse Modes in Hindi Stories (2020.lrec-1)

Copied to clipboard

Challenge: Using a new corpus of sentences from Hindi short stories, we analyze the annotations for five different discourse modes argumentative, narrative, descriptive, dialogic and informative.
Approach: They propose to annotate sentences from Hindi short stories for five different discourse modes argumentative, narrative, descriptive, dialogic and informative.
Outcome: The proposed corpus has a high inter-annotator agreement (0.87 k-alpha) and is able to capture the nuances of the embedded discourse structures.
Help! Need Advice on Identifying Advice (2020.emnlp-main)

Copied to clipboard

Challenge: Pre-trained systems are able to capture advice better than rule-based systems, but advice identification is challenging.
Approach: They analyze a dataset of advice posts on two reddit forums and annotate whether they contain advice.
Outcome: The proposed models show that pre-trained models capture advice better than rule-based systems, but advice identification is challenging.
Discourse Analysis via Questions and Answers: Parsing Dependency Structures of Questions Under Discussion (2023.findings-acl)

Copied to clipboard

Challenge: Existing discourse formalisms require large taxonomies of discourse relations to be accurate.
Approach: They propose a linguistic framework for discourse analysis using questions under discussion . they propose qUD parser that derives a dependency structure of questions over full documents .
Outcome: The proposed model is trained on a large, crowdsourced question-answering dataset.
Unsupervised Extractive Summarization of Emotion Triggers (2023.acl-long)

Copied to clipboard

Challenge: Recent approaches trained supervised models to detect emotions and explain emotion triggers via abstractive summarization, but this can block necessary responses.
Approach: They propose to augment an abstractive dataset with extractive triggers and develop unsupervised models that can jointly detect emotions and summarize their triggers.
Outcome: The proposed model outperforms existing models and is based on a COVID-19 crisis dataset.
How Do We Answer Complex Questions: Discourse Structure of Long-form Answers (2022.acl-long)

Copied to clipboard

Challenge: Recent work explored long-form answers, where answers are free-form texts consisting of multiple sentences.
Approach: They develop an ontology of six sentence-level functional roles for long-form answers . they annotate 3.9k sentences in 640 answer paragraphs and train a strong classifier .
Outcome: The proposed model-generated answers agree less with model-driven answers than human-written answers.
Did they answer? Subjective acts and intents in conversational discourse (2021.naacl-main)

Copied to clipboard

Challenge: Discourse signals are often implicit, leaving it up to the interpreter to draw inferences . current discourse data and frameworks ignore the social aspect, expecting only a single ground truth . elisa f. and her team present a dataset with multiple and subjective interpretations of English conversation .
Approach: They present a first discourse dataset with multiple and subjective interpretations of English conversation . they show disagreements are nuanced and require a deeper understanding of contextual factors .
Outcome: The proposed dataset shows disagreements are nuanced and require deeper understanding of contextual factors.
Counterfactual Probing for the Influence of Affect and Specificity on Intergroup Bias (2023.findings-acl)

Copied to clipboard

Challenge: Existing work on bias in NLP only considers negative or pejorative language use.
Approach: They propose a revised framing of bias in terms of intergroup social context and its effects on language output.
Outcome: The proposed framework is based on a model of intergroup relationships in English language tweets.
Which questions should I answer? Salience Prediction of Inquisitive Questions (2024.emnlp-main)

Copied to clipboard

Challenge: Recent work in NLP has taken advantage of question generation capabilities of LLMs to enhance a wide range of applications.
Approach: They propose a salience predictor for inquisitive questions that is instruction-tuned . they show that highly salient questions are empirically more likely to be answered in the same article .
Outcome: The proposed model is based on linguist-annotated salience scores of 1,766 questions . it shows that answering salient questions improves comprehension of the text .
Evaluating Discourse in Structured Text Representations (P19-1)

Copied to clipboard

Challenge: Discourse structure is integral to understanding a text and is useful in many NLP tasks.
Approach: They propose a structured attention mechanism for text classification that derives a tree over a text, akin to an RST discourse tree.
Outcome: The proposed model improves performance on multiple discourse-relevant tasks and datasets and ablation studies show it does little to capture discourse structure.
Behavioral Analysis of Information Salience in Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models excel at text summarization, but the exact notion of salience remains unclear.
Approach: They propose a framework to derive and investigate information salience in Large Language Models (LLMs) using length-controlled summarization as a behavioral probe into the content selection process.
Outcome: The proposed framework derives a proxy for how models prioritize information in large language models.
Text Simplification of College Admissions Instructions: A Professionally Simplified and Verified Corpus (2022.coling-1)

Copied to clipboard

Challenge: a dataset of 112 admissions instructions is used to simplify the language used by higher education institutions to communicate with prospective students.
Approach: They propose to simplify admissions instructions by professionally simplifying them and comparing them to a dataset of 112 admissions documents.
Outcome: The proposed dataset includes 112 admissions instructions from higher education institutions across the US.
Using Developer Discussions to Guide Fixing Bugs in Software (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent work shows that natural language context is useful in guiding bug-fixing models, but requires prompting developers to provide this context.
Approach: They propose to use bug report discussions to prompt developers to provide natural language context for bug-fixing models.
Outcome: The proposed approach reduces the need for additional information from developers.
Wugnectives: Novel Entity Inferences of Language Models from Discourse Connectives (2026.eacl-long)

Copied to clipboard

Challenge: Using context + knowledge of discourse connectives to make predictions about discourse connective .
Approach: They present a dataset of 8,880 stimuli that evaluates LMs’ inferences about novel entities in contexts where connectives link the entities to particular attributes.
Outcome: The proposed dataset evaluates LMs’ inferences about new entities in contexts where connectives link the entities to particular attributes.
Is It JUST Semantics? A Case Study of Discourse Particle Understanding in LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Discourse particles are crucial elements that subtly shape the meaning of text.
Approach: They examine the capacity of linguists to distinguish fine-grained senses of English *just* . they find that they struggle to fully capture more subtle nuances of discourse particles .
Outcome: The study shows that linguists struggle to capture subtle nuances of discourse particles.
The Role of Context and Uncertainty in Shallow Discourse Parsing (2022.coling-1)

Copied to clipboard

Challenge: Discourse parsing has proven to be useful for a number of NLP tasks that require complex reasoning.
Approach: They hypothesize that context plays an important role in accurate human annotation and add uncertainty measures can improve model accuracy and calibration.
Outcome: The proposed model can be better calibrated by adding uncertainty measures to models with better accuracy and calibration.
SNaC: Coherence Error Detection for Narrative Summarization (2022.emnlp-main)

Copied to clipboard

Challenge: SNaC framework is used to evaluate long summaries, but it fails to identify gaps in coherence . nallapati and colleagues have developed a framework for fine-grained annotations of long summarizations .
Approach: They propose a narrative coherence evaluation framework for fine-grained annotations of long summaries that can be used to evaluate coherent narratives.
Outcome: The proposed framework can support future work in document summarization and coherence evaluation, the authors show .
Sarcasm Detection in a Disaster Context (2024.lrec-main)

Copied to clipboard

Challenge: During natural disasters, people often use social media platforms to express contempt or sarcasm . despite being widely researched as an NLP task, sarkasmatic detection has not been explored in a specific context .
Approach: They propose a dataset of 15,000 tweets annotated for intended sarcasm . they propose sarkasmatic detection using pre-trained language models .
Outcome: The proposed model can obtain as much as 0.70 F1 on the dataset.
Language Models (Mostly) Do Not Consider Emotion Triggers When Predicting Emotion (2024.naacl-short)

Copied to clipboard

Challenge: Existing work has sought to identify what triggers or causes a particular emotion, but the relationship between those triggers and the prediction of emotion detection models is little understood.
Approach: They propose a dataset to evaluate the ability of large language models to identify emotion triggers . they compare features considered important for emotion prediction models to those considered less salient .
Outcome: The proposed dataset compares large language models and fine-tuned models on social media posts . it shows that emotion triggers are not considered salient features for emotion prediction models .
Detecting Perceived Emotions in Hurricane Disasters (2020.acl-main)

Copied to clipboard

Challenge: Existing methods for emotion detection are limited in disaster-centric domains due to distributional shifts.
Approach: They propose to use a Twitter emotion dataset to analyze emotions in natural disasters . they propose to apply classification tasks to discriminate between coarse-grained emotions .
Outcome: The proposed model achieves only 68% accuracy after pre-training with unlabeled Twitter data.
Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence (2026.findings-acl)

Copied to clipboard

Challenge: Existing models are overwhelmingly accurate when presented with counterfactual medical evidence . prior work explored conflicts between context and LLM parametric knowledge in the general domain .
Approach: They construct a counterfactual medical QA dataset that requires models to answer clinical comparison questions with evidence from randomized controlled trials.
Outcome: The proposed model overemphasizes the latter, and the model overestimates the latter.
FALTE: A Toolkit for Fine-grained Annotation for Long Text Evaluation (2022.emnlp-demos)

Copied to clipboard

Challenge: Existing tools to evaluate long text outputs are lacking in the field of NLP . human rating and error analysis remains a crucial component for any evaluation of long text generation.
Approach: They propose a web-based toolkit to collect fine-grained error annotations for long texts . they use a taxonomy to identify errors and assign them to text spans .
Outcome: The proposed tool can be used to evaluate the coherence of long generated summaries.
Political Ideology and Polarization: A Multi-dimensional Approach (2022.naacl-main)

Copied to clipboard

Challenge: Recent research has made great strides towards understanding the ideological bias (i.e., stance) of news media along the left-right spectrum.
Approach: They propose a novel approach for the study of ideology based on its left or right positions on the issue being discussed.
Outcome: The proposed method allows for the quantitative and temporal measurement and analysis of polarization as a multidimensional ideological distance.
A Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical Literature (P18-1)

Copied to clipboard

Challenge: In 2015 alone, about 100 manuscripts describing randomized controlled trials for medical interventions were published every day.
Approach: They propose a corpus of 5,000 medical articles annotated with demarcations of text spans that describe the Patient population enrolled, the Interventions studied and to what they were Compared, and the Outcomes measured.
Outcome: The proposed corpus includes 5,000 medical articles describing clinical randomized controlled trials.
Learning to Describe Solutions for Bug Reports Based on Developer Discussions (2022.findings-acl)

Copied to clipboard

Challenge: Software bugs in open-source projects are reported through issue tracking systems like GitHub Issues.
Approach: They propose a method for generating a natural language description of a bug by synthesizing relevant content within the discussion.
Outcome: The proposed system generates a natural language description of the solution by synthesizing relevant content within the discussion.
Emotion analysis and detection during COVID-19 (2022.lrec-1)

Copied to clipboard

Challenge: 3,000 English tweets labeled with emotions are used to predict emotions during crises . authors propose semi-supervised learning to bridge this gap .
Approach: They propose to use a dataset of 3,000 English tweets labeled with emotions . they propose semi-supervised learning to bridge this gap by analyzing unlabeled data .
Outcome: The proposed model can be used to predict emotions in the context of COVID-19 . the proposed model performs better than other models using unlabeled data .
Do *they* mean ‘us’? Interpreting Referring Expression variation under Intergroup Bias (2024.findings-emnlp)

Copied to clipboard

Challenge: We model intergroup bias as a tagging task on English sports comments from forums dedicated to fandom for NFL teams . linguistic descriptions of win probability are used for large-scale analysis of intergroup variation .
Approach: They propose to model intergroup bias as a tagging task on NFL fan comments . they use linguistic models to model the bias and use them to generate large-scale annotations .
Outcome: The proposed model can reveal unobserved variations in the form of referents across win probabilities.
Impact of Evaluation Methodologies on Code Summarization (2022.acl-long)

Copied to clipboard

Challenge: Existing evaluation methodologies for code summarization tasks do not consider timestamps of code and comments.
Approach: They propose a time-segmented evaluation methodology for code summarization that considers timestamps of code and comments during evaluation.
Outcome: The proposed evaluation methodology compares with other evaluation methodologies that have been widely used.
Discourse Comprehension: A Question Answering Framework to Represent Sentence Connections (2022.emnlp-main)

Copied to clipboard

Challenge: Existing systems for text comprehension are inadequate for more holistic comprehension of a discourse.
Approach: They propose a new paradigm that captures both discourse and semantic links between sentences in the form of free-form, open-ended questions.
Outcome: The proposed model captures discourse and semantic links between sentences in the form of free-form, open-ended questions.
Improving the Distributional Alignment of LLMs using Supervision (2026.acl-long)

Copied to clipboard

Challenge: Existing work to evaluate LLMs' alignment with human values and opinions has a key shortcoming.
Approach: They propose to add supervision to LLMs to improve alignment with diverse populations . they find that supervision improves alignment across public health, public opinion, values and beliefs .
Outcome: The proposed method improves the alignment of LLMs with diverse populations on subjective questions.
Decide less, communicate more: On the construct validity of end-to-end fact-checking in medicine (2026.findings-acl)

Copied to clipboard

Challenge: Evidence-based medicine connects to every individual, yet the nature of it is highly technical . e-fact-checking systems that connect to medical decisions are largely unused . we examine how clinical experts verify real claims from social media .
Approach: They propose that fact-checking should be approached as an interactive communication problem . they argue that social media and AI have made medical knowledge accessible .
Outcome: The proposed method is based on the work of a clinical expert on social media . it reveals that the method is difficult to connect claims to clinical trials .
FactPICO: Factuality Evaluation for Plain Language Summarization of Medical Evidence (2024.acl-long)

Copied to clipboard

Challenge: FactPICO is a factuality benchmark for plain language summarization of medical texts describing randomized controlled trials . existing metrics for factual summarizing medical evidence are poorly correlated with expert judgments on the instance level.
Approach: They propose a factuality benchmark for plain language summarization of medical texts . they assess factuality of critical elements of RCTs in those summaries .
Outcome: The proposed benchmark assesses the factuality of medical summaries using LLMs . the summary summators are based on 345 plain language summaires with fine-grained evaluation .
Learning to Refine with Fine-Grained Natural Language Feedback (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent work has explored the capability of large language models to identify and correct errors in LLM-generated responses.
Approach: They propose to combine refinement with feedback into three distinct competencies . step 1: Detect, Critique, Refine gives a fine-grained feedback about errors .
Outcome: The proposed method outperforms existing refinement approaches and models not fine-tuned for factuality critiquing.
Evaluating Factuality in Text Simplification (2022.acl-long)

Copied to clipboard

Challenge: Automated simplification models aim to make input texts more readable without altering their meaning.
Approach: They propose a taxonomy of errors that are used to analyze simplification models . they propose to use simplification methods to make input texts more readable .
Outcome: The proposed models introduce errors that are not captured by existing evaluation metrics.
Expressively vulgar: The socio-dynamics of vulgarity and its effects on sentiment analysis in social media (C18-1)

Copied to clipboard

Challenge: Vulgarity is a common linguistic expression and is used to perform several linguistic functions.
Approach: They analyze vulgarity using tweets from users with known demographics and sentiment ratings for vulgar tweets to study sentiment analysis performance.
Outcome: The proposed model can boost sentiment analysis performance by analyzing vulgar tweets and tweet sentiment ratings.
Inquisitive Question Generation for High Level Text Comprehension (2020.emnlp-main)

Copied to clipboard

Challenge: Existing data-driven questions generate questions that fill gaps in knowledge . a dataset of 19K questions is used to generate meaningful questions .
Approach: They propose a dataset of 19K questions that are elicited while a person is reading a document.
Outcome: The proposed model generates reasonable questions, but the task is challenging.
Why Do You Feel This Way? Summarizing Triggers of Emotions in Social Media Posts (2022.emnlp-main)

Copied to clipboard

Challenge: Large-scale crises such as the COVID-19 pandemic cause emotional turmoil worldwide.
Approach: They propose a method to jointly detect emotions and summarize emotion triggers in social media posts related to COVID-19.
Outcome: The proposed method can detect emotions and summarize emotions in long social media posts.
How people talk about each other: Modeling Generalized Intergroup Bias and Emotion (2023.eacl-main)

Copied to clipboard

Challenge: Current studies of bias in NLP rely on identifying (unwanted or negative) bias towards a specific demographic group, but this is not always practical.
Approach: They extrapolate a notion of bias from social science literature to predict interpersonal group relationship (IGR) using interpersonal emotions as an anchor.
Outcome: The proposed model predicts the interpersonal group relationship (IGR) using interpersonal emotions as an anchor.
Elaborative Simplification as Implicit Questions Under Discussion (2023.emnlp-main)

Copied to clipboard

Challenge: Automated text simplification is often thought of as a monolingual translation task . this view fails to account for elaborative simplification, where new information is added into the simplified text.
Approach: They propose to view elaborative simplification through the lens of the Question Under Discussion framework . they propose to model 1.3K elongations accompanied by implicit QUDs to investigate what writers elaborate upon .
Outcome: The proposed framework provides a robust way to investigate what writers elaborate upon, how they elaborate, and how elaborations fit into the discourse context.
ProtoTEx: Explaining Model Decisions with Prototype Tensors (2022.acl-long)

Copied to clipboard

Challenge: Neural models for NLP have yielded significant gains in predictive accuracy across tasks.
Approach: They propose a white-box NLP classification architecture based on prototype networks . they propose an interleaved training algorithm that faithfully explains model decisions .
Outcome: The proposed model matches BART-large and exceeds BERTlarge on propaganda detection tasks.
Adaptive Ensembling: Unsupervised Domain Adaptation for Political Document Analysis (D19-1)

Copied to clipboard

Challenge: a new study examines the use of labeled and unlabeled corpora in political science research . large corporata often contain documents of a certain subject or type, but they are often unlabed . a recent study found that labeles with pertinent documents stem from a single source .
Approach: They propose an unsupervised domain adaptation framework that uses a text classification model and time-aware training to ensure it works well with diachronic corpora.
Outcome: The proposed framework outperforms benchmarks on an expert-annotated dataset and is more stable and learns better representations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations